Papers with domain specificity

6 papers
Can Automatic Post-Editing Improve NMT? (2020.emnlp-main)

Copied to clipboard

Challenge: APE has been successful with statistical machine translation systems but has not been as successful over neural machine translation (NMT) systems.
Approach: They propose to train neural APE models on a corpus of human post-edits of NMT and compile a larger corpus to test their hypothesis.
Outcome: The proposed model can improve a strong in-domain NMT system, challenging the current understanding in the field.
What Makes a Good Query? Measuring the Impact of Human-Confusing Linguistic Features on LLM Performance (2026.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are often treated as defects of the model or its decoding strategy.
Approach: They construct a 22-dimension query feature vector covering clause complexity, lexical rarity, anaphora, negation, answerability, and intention grounding.
Outcome: The proposed model covers clause complexity, lexical rarity, anaphora, negation, answerability, and intention grounding, all known to affect human comprehension.
LlamaLens: Specialized Multilingual LLM for Analyzing News and Social Media Content (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable success as general-purpose task solvers across various fields.
Approach: They propose to develop a specialized LLM for analyzing news and social media content in a multilingual context.
Outcome: The proposed model outperforms the current state-of-the-art on 23 testing sets and achieves comparable performance on 8 sets.
Iterative Constrained Back-Translation for Unsupervised Domain Adaptation of Machine Translation (2022.coling-1)

Copied to clipboard

Challenge: Existing back-translation methods focus on in-domain lexical knowledge, which may lead to poor translation of unseen in- domain words.
Approach: They propose an iterative constrained back-translation method to incorporate in-domain lexical knowledge into synthetic parallel data from BT.
Outcome: The proposed method improves the BLEU score by up to 3.08 on four domains.
Multilingual Data Filtering using Synthetic Data from Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have shown that effective filters can be created by utilising Large Language Models to synthetically label data, which is then used to train smaller neural models for filtering purposes.
Approach: They extend this approach to languages beyond English to train neural models for filtering purposes.
Outcome: The proposed approach is effective at filtering parallel text for translation quality and filtering for domain specificity.
TechniqueRAG: Retrieval Augmented Generation for Adversarial Technique Annotation in Cyber Threat Intelligence Text (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for identifying adversarial techniques in security texts face a trade-off: generic models with limited domain precision or resource-intensive pipelines.
Approach: They propose a domain-specific retrieval-augmented generation framework that integrates off-the-shelf retrievers, instruction-tuned LLMs, and minimal text–technique pairs.
Outcome: The proposed framework improves retrieval quality and domain specificity without extensive optimizations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations